NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Orchestrating a DNN training job using an iScheduler Framework: a use case.

Vallabhajosyula, ManikyaSwathi; Budhya, Sandeep Satish; Jain, Akanksha; Baig, Maaz; Ramnath, Rajiv (July 2024, In Practice and Experience in Advanced Research Computing (PEARC '24))

Full Text Available
Effective Mimicry of Belady’s MIN Policy

https://doi.org/10.1109/HPCA53966.2022.00048

Shah, Ishan; Jain, Akanksha; Lin, Calvin (April 2022, International Symposium on High-Performance Computer Architecture)

The past decade has seen the rise of highly successful cache replacement policies that are based on binary prediction. For example, the Hawkeye policy learns whether lines loaded by a given PC are Cache Friendly (likely to remain in the cache if Belady’s MIN policy had been used) or Cache Averse (likely to be evicted by Belady’s MIN policy). In this paper, we instead present a cache replacement policy that is based on multiclass prediction, which allows it to directly mimic Belady’s MIN policy in a surprisingly simple and effective way. Our policy uses a PC-based predictor to learn each cache line’s reuse distance; it then evicts lines based on their predicted time of reuse. We show that our use of multiclass prediction is more effective than binary prediction because it allows for a finer-grained ordering of cache lines during eviction and because it is more robust to prediction errors.Our empirical results show that our new policy, which we refer to as Mockingjay, outperforms the previous state-of-the-art on both single-core and multi-core platforms and both with and without a prefetcher. For example, with no prefetcher, on a mix of 100 multi-core workloads from the SPEC 2006, SPEC 2017, and GAP benchmark suites, Mockingjay sees an average improvement over LRU of 15.2%, compared to 7.6% for SHiP and 12.9% for Hawkeye. On a single-core platform, Mockingjay’s improvement over LRU is 5.7%, which approaches the 6.0% improvement of Belady MIN’s unrealizable policy. On a single-core platform (with a prefetcher) running the high-MPKI CVP workloads, Mockingjay’s improvement over LRU is 20.1%, compared to 13.4% for Hawkeye.
more » « less
Full Text Available
Practical Temporal Prefetching With Compressed On-Chip Metadata

https://doi.org/10.1109/TC.2021.3065909

Wu, Hao; Nathella, Krishnendra; Pabst, Matthew; Sunwoo, Dam; Jain, Akanksha; Lin, Calvin (October 2021, IEEE Transactions on Computers)

Temporal prefetchers are powerful because they can prefetch irregular sequences of memory accesses, but temporal prefetchers are commercially infeasible because they store large amounts of metadata in DRAM. This paper presents Triage, the first temporal data prefetcher that does not require off-chip metadata. Triage builds on two insights: (1) Metadata are not equally useful, so the less useful metadata need not be saved, and (2) for irregular workloads, it is more profitable to use portions of the LLC to store metadata than data. We also introduce novel schemes to identify useful metadata, to compress metadata, and to determine the fraction of the LLC to dedicate for metadata.
more » « less
Full Text Available
A Hierarchical Neural Model of Data Prefetching

https://doi.org/10.1145/3445814.3446752

Shi, Zhan; Jain, Akanksha; Swersky, Kevin; Hashemi, Milad; Ranganathan, Parthasarathy; Lin, Calvin (April 2021, nternational Conference on Architectural Support for Programming Languages and Operating Systems)

This paper presents Voyager, a novel neural network for data prefetching. Unlike previous neural models for prefetching, which are limited to learning delta correlations, our model can also learn address correlations, which are important for prefetching irregular sequences of memory accesses. The key to our solution is its hierarchical structure that separates addresses into pages and offsets and that introduces a mechanism for learning important relations among pages and offsets. Voyager provides significant prediction benefits over current data prefetchers. For a set of irregular programs from the SPEC 2006 and GAP benchmark suites, Voyager sees an average IPC improvement of 41.6% over a system with no prefetcher, compared with 21.7% and 28.2%, respectively, for idealized Domino and ISB prefetchers. We also find that for two commercial workloads for which current data prefetchers see very little benefit, Voyager dramatically improves both accuracy and coverage. At present, slow training and prediction preclude neural models from being practically used in hardware, but Voyager’s overheads are significantly lower—in every dimension—than those of previous neural models. For example, computation cost is reduced by 15- 20×, and storage overhead is reduced by 110-200×. Thus, Voyager represents a significant step towards a practical neural prefetcher.
more » « less
Full Text Available
Cache Replacement Policies

https://doi.org/10.2200/S00922ED1V01Y201905CAC047

Jain, Akanksha; Lin, Calvin (June 2019, Synthesis Lectures on Computer Architecture)

Full Text Available
Efficient metadata management for irregular data prefetching

https://doi.org/10.1145/3307650.3322225

Wu, Hao; Nathella, Krishnendra; Sunwoo, Dam; Jain, Akanksha; Lin, Calvin (June 2019, Efficient metadata management for irregular data prefetching)

Temporal prefetchers have the potential to prefetch arbitrary memory access patterns, but they require large amounts of metadata that must typically be stored in DRAM. In 2013, the Irregular Stream Buffer (ISB), showed how this metadata could be cached on chip and managed implicitly by synchronizing its contents with that of the TLB. This paper reveals the inefficiency of that approach and presents a new metadata management scheme that uses a simple metadata prefetcher to feed the metadata cache. The result is the Managed ISB (MISB), a temporal prefetcher that significantly advances the state-of-the-art in terms of both traffic overhead and IPC. Using a highly accurate proprietary simulator for single-core workloads, and using the ChampSim simulator for multi-core workloads, we evaluate MISB on programs from the SPEC CPU 2006 and CloudSuite benchmarks suites. Our results show that for single-core workloads, MISB improves performance by 22.7%, compared to 10.6% for an idealized STMS and 4.5% for a realistic ISB. MISB also significantly reduces off-chip traffic; for SPEC, MISB's traffic overhead of 70% is roughly one fifth of STMS's (342%) and one sixth of ISB's (411%). On 4-core multi-programmed workloads, MISB improves performance by 27.5%, compared to 13.6% for idealized STMS. For CloudSuite, MISB improves performance by 12.8% (vs. 6.0% for idealized STMS), while achieving a traffic reduction of 7 × (83.5% for MISB vs. 572.3% for STMS).
more » « less
Full Text Available
Combining Branch History and Value History For Improved Value Prediction

Sakhuja, Chirag; Subramanian, Anjana; Joshi, Pawan; Jain, Akanksha; and Lin, Calvin (January 2019, Second Championship Value Prediction)
null (Ed.)
State-of-the-art value predictors either use control-flow context or data context to predict values. Predictors based on control-flow context use branch histories to remember past values, but these predictors require lengthy histories to predict anything other than constant and strided values. Predictors that use data context---also known as Finite Context Method (FCM) predictors---use a history of past values to predict a broader class of values, but such predictors achieve low coverage due to long training times, and they can become complex due to speculative value histories. We observe that the combination of branch and value history provides better predictability than the use of each history separately because it can predict values in control-dependent sequences of values. Furthermore, the combination improves training time by enabling accurate predictions to be made with shorter history, and it simplifies the hardware design by removing the need for speculative value histories. Based on these observations, we propose a new unlimited budget value predictor, Heterogeneous-Context Value Predictor (HCVP), that when hybridized with E-Stride, achieves a geometric mean IPC of 3.88 on the 135 public traces, as compared to 3.81 for the current leader of the Championship Value Prediction.
more » « less
Full Text Available

Search for: All records